This notebook will examine some characterisitics of the "housing-header.txt" dataset, a dataset of housing data in Boston neighborhoods. For reference:

  1. CRIM: per capita crime rate by town
  2. ZN: proportion of residential land zoned for lots over 25,000 sq.ft .
  3. INDUS: proportion of non retail business acres per town
  4. CHAS: Charles River dummy variable (= 1 if tract bounds river; 0 otherwise)
  5. NOX: nitric oxides concentration (parts per 10 million)
  6. RM: average number of rooms per dwelling
  7. AGE: proportion of owner occupied units built prior to 1940
  8. DIS: weighted distances to five Boston employment centers
  9. RAD: index of accessibility to radial highways
  10. TAX: full value property tax rate per 10,000
  11. PTRATIO: pupil teacher ratio by town
  12. B: 1000(Bk 0.63)^2 where Bk is the proportion of blacks by town
  13. LSTAT: percent lower status of the population
  14. MEDV: Median value of owner occupied homes in 1000's

First we must load the dataset:

In [1]:
housing<-read.table("housing.header.txt", header=TRUE, sep=",")
colnames(housing)<-c("Crim", "Zn", "Indus", "Chas", "Nox", "Rm", "Age", "Dis", "Rad", "Tax", "Ptratio", "B", "Lstat", "Medv")
head(housing)
CrimZnIndusChasNoxRmAgeDisRadTaxPtratioBLstatMedv
0.0063218 2.31 0 0.538 6.575 65.2 4.0900 1 296 15.3 396.90 4.98 24.0
0.02731 0 7.07 0 0.469 6.421 78.9 4.9671 2 242 17.8 396.90 9.14 21.6
0.02729 0 7.07 0 0.469 7.185 61.1 4.9671 2 242 17.8 392.83 4.03 34.7
0.03237 0 2.18 0 0.458 6.998 45.8 6.0622 3 222 18.7 394.63 2.94 33.4
0.06905 0 2.18 0 0.458 7.147 54.2 6.0622 3 222 18.7 396.90 5.33 36.2
0.02985 0 2.18 0 0.458 6.430 58.7 6.0622 3 222 18.7 394.12 5.21 28.7

We then plot all samples with respect to the crime rate:

In [2]:
attach(housing)
plot(Crim)

We then continue with our example using crime rate: plotting its density as well as a histogram:

In [3]:
par(mfrow=c(1,2))
hist(Crim)
title ("Histogram of Crim")
plot(density(Crim, na.rm=TRUE))

We then plot crime rate, number of rooms, age of the home, and property tax rate with respect to median house value, in order to examine the correlation:

In [4]:
par(mfrow=c(1,4))
plot(Crim, Medv)
title("Crim-Medv")
plot(Rm, Medv)
title("Rm-Medv")
plot(Age, Medv)
title("Age-Medv")
plot(Tax, Medv)
title("Tax-Medv")

From the scatterplots, we can observe that:

  • Crim and Medv appear to have a negative correlation
  • Rm and Medv appear to have a positive correlation
  • Age and Medv appear to have a negative correlation
  • Tax and Medv appear to have a negative correlation

We can then report the pairwise correlation between every two variables using a level plot:

In [5]:
cv<-cor(housing)
library(lattice)
levelplot(cv)

Observe the level plot above. The following variables are negatively correlated with Medv:

  • Lstat
  • Ptratio
  • Tax
  • Rad
  • Age
  • Nox
  • Indus
  • Crim

The following variables have little or no correlation with Medv:

  • B
  • Dis
  • Chas
  • Zn

The following variables are positively correlated with Medv:

  • Medv itself (correlation of 1 expected)
  • Rm

We also demonstrate this same correlation with scatterplots:

In [6]:
par(mfrow=c(2, 7))
plot(Crim, Medv)
title("Crim-Medv")
plot(Zn, Medv)
title("Zn-Medv")
plot(Indus,Medv)
title("Indus-Medv")
plot(Chas, Medv)
title("Chas-Medv")
plot(Nox,Medv)
title("Nox-Medv")
plot(Rm, Medv)
title("Rm-Medv")
plot(Age, Medv)
title("Age-Medv")
plot(Dis, Medv)
title("Dis-Medv")
plot(Rad, Medv)
title("Rad-Medv")
plot(Tax, Medv)
title("Tax-Medv")
plot(Ptratio, Medv)
title("Ptratio-Medv")
plot(B, Medv)
title("B-Medv")
plot(Lstat, Medv)
title("Lstat-Medv")

The attributes that are posiitvely correlated with Medv will have a generally increasing trend on the scatterplot (x increases as y increases). The attributes that are independent of Medv will have no easily observable trend. The attributes that are negatvely correlated with Medv will have a generally decreasing trend on the scatterplot (x decreases as y increases)

In [ ]: